Papers with adversarial annotator

1 papers
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have significantly enhanced the capabilities in natural language processing.
Approach: They propose a method to poison large language models by using annotators to rank a set of collected responses to generate longer tokens.
Outcome: The proposed method can generate longer tokens without harming the original safety alignment performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations